Papers with Stable Diffusion XL
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation (2024.emnlp-main)
Copied to clipboard
Raphael Tang, Crystina Zhang, Lixinyu Xu, Yao Lu, Wenyan Li, Pontus Stenetorp, Jimmy Lin, Ferhan Ture
| Challenge: | Current diffusion models do not cover recent models, thus we curate three test sets for evaluation. |
| Approach: | They propose a human-calibrated measure of variability in a set of images bootstrapped from existing image-pair perceptual distances. |
| Outcome: | The proposed model outperforms nine baselines by 18 points in accuracy and matches graded human judgements 78% of the time. |
Minimal, Local, and Robust: Embedding-Only Edits for Implicit Bias in T2I Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | EmbEdit is a text-to-image editing method that only fine-tunes the word token embedding (WTE) of the target object. |
| Approach: | They propose a method to edit implicit assumptions and priors in text-to-image models without affecting unrelated objects or degrading overall performance. |
| Outcome: | The proposed method outperforms previous methods in various models, tasks, and editing scenarios. |